Skip to content

docs(native-surfaces): record 2.1.284 overlap verdicts and Boundary sections - #5387

Merged
kyle-sexton merged 25 commits into
mainfrom
feat/native-surface-verdicts-2.1.284
Sep 29, 2026
Merged

kyle-sexton merged 25 commits into
mainfrom
feat/native-surface-verdicts-2.1.284

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

No related issue: follow-up to #5371, recording verdicts for native Claude Code surfaces the 2.1.284 overlap detector found.

Summary

Claude Code 2.1.284 ships native surfaces that overlap this marketplace's skills: bundled commit and pr, /autofix-pr, /subtask, /fork, /background, /recap, the deep-research workflow, /doctor prompt-audit, /verify, /batch, and others. Nothing told the model when to prefer the native surface, or that a user-only command exists to offer the person.

Stacked on #5371; merge that first.

Fix

  • docs/native-surfaces/records.json: 25 new rows (22 complementary, 3 defer), and every existing extraction row re-derived against the 2.1.284 extraction. No existing verdict changed. Marker drift was corrected where no verdict depended on it: claude-api ×3 gated, design-sync not hidden, and skill-doctor and export ×3 model-invocation-disabled. View regenerated (47 rows).
  • Rulings: made 2026-09-29 by operator direction on the orchestrator's recommendation. Each reason says so; this review is the human gate on them.
  • ## Boundary sections: one per non-defer row, each with a four-part reference file in the same skill. The sections are presence-gated ("when the bundled pr skill resolves in this session…"), name the native surface in a code span, and copy none of its behavior.
    • A user-only surface (integration: suggest) is offered to the person: "you can run /autofix-pr instead of or alongside this".
    • A model-invocable surface gets a routing split.
  • Plugins touched (patch bump and CHANGELOG entry each): source-control, claude-config, verification, implementation, discovery, claude-ops, context-budget, session-flow, planning, prototype, debugging.
  • discovery:research-deep no longer claims it can dispatch the bundled deep-research workflow, because that workflow has model invocation disabled. It now offers the workflow to the person.
  • Frontmatter descriptions are not edited. Description phrases follow one plugin per PR under the sweep contract in audit-native-overlap/SKILL.md.

Verification

  • overlap.py generate --check: in sync (47 rows).
  • overlap.py self-check: degraded, with the 2 documented advisories only: older recorded versions in rows left for review, and no --upstream-sha.
  • test_overlap.py: 134 tests OK.
  • check-changed-skills.sh main: 20 skills, 0 failed.
  • check-changelog-parity.sh: --check, --check-bump main, --check-order and --check-preserved main all exit 0.

Related

🤖 Generated with Claude Code

kyle-sexton and others added 21 commits September 29, 2026 13:53
The brace reader treated template-literal ${...} substitutions as text, so
quotes inside a regex in a substitution desynchronized it and one brace pair
swallowed 21 MB of the bundle; 15 of 152 commands resolved. Substitutions are
now tokenized as code. Also resolves registerSlidesSkill, literal-table skill
rosters, and constant-named commands; tightens registrar and registration-token
matching instead of widening thresholds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
detect now scores every native surface (builtin commands, bundled skills,
plugin-backed built-ins, and bundled workflows when the inventory carries
that lane) against every repo skill and agent from name and description
tokens, and emits pairs over a threshold, top-k per surface, as
origin "discovered" beside the seeded pairs. Pairs already in the store are
listed as existing with their verdict; seeds absorb their discovered twin.

Each candidate carries invocable_by from model_invocable/user_invocable
(older inventories degrade to unknown) and a recommended_integration label;
model_invocable false sets the model-invocation-disabled marker the store's
suggest-only rule reads. bundled-workflow joins the provenance classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven pairs whose overlap is conceptual rather than lexical (recap, fork,
subtask, batch, explain-usage x2, fewer-permission-prompts) score below the
discovery cut against Claude Code 2.1.284, so they join the seeded pairs.
Seeded candidates now report their lexical score even below the cut.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The report structure now shows each candidate's origin, score, invocable_by
and recommended integration, and the detection posture states how discovery
scores and where seeds still earn their place. The plugin_backed lane and
code-review alias gotchas are re-verified against Claude Code 2.1.284.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cs cross-check

The inventory now emits argument_hint and description resolved from
getters, constants, function references, and concatenations; user_invocable
and model_invocable on every command and bundled skill, null when the
bundle decides at runtime; a bundled_workflows lane with a deep-research
canary; and a --docs mode that classifies each name against the commands
page and attaches changelog history as a labeled heuristic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s cross-check

SKILL.md gains the Invocable-by marker, the bundled workflows and docs
cross-check report sections, the --docs flags, and a Next pointer to the
native-overlap audit. extraction.md covers the field resolver, the
invocability rules, the workflow push-site registrar, and the docs lane.
Verification records re-checked against Claude Code 2.1.284 and the
2026-09-29 docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rfaces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…faces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rfaces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rfaces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aces

Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Records how each component relates to the Claude Code 2.1.284 native surfaces the operator ruled on, with four-part records in the skill's own reference file. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds 21 complementary rows (6 route, 15 suggest) and 3 defer rows ruled 2026-09-29 by operator direction, refreshes 8 extraction rows whose surface, class and markers still match the 2.1.284 extraction, and marks the existing skill-doctor and playground Boundary sections baked. The debug -> debugging:debug pair is not recorded: the bundled skill disables model invocation, which contradicts the requested route integration.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The bundled skill debugs Claude Code itself and is reserved for the person to run; the model offers /debug when the problem is Claude Code rather than the user's application. Patch bump and CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds debug -> debugging:debug (complementary, suggest). Adds gated to the three claude-api rows, drops hidden from design-sync, adds model-invocation-disabled to skill-doctor and the three export rows, all per the 2.1.284 extraction. Drops stemmed detect tokens from evidence lines so typos passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ences

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Base automatically changed from fix/claude-ops-native-inventory-2.1.284 to main September 29, 2026 21:58
…erdicts-2.1.284

# Conflicts:
#	docs/conventions/native-references/CHANGELOG.md
#	docs/native-surfaces.md
#	plugins/claude-config/CHANGELOG.md
#	plugins/claude-ops/.claude-plugin/plugin.json
#	plugins/claude-ops/CHANGELOG.md
#	plugins/claude-ops/skills/audit-native-overlap/SKILL.md
#	plugins/claude-ops/skills/audit-native-overlap/scripts/discover.py
#	plugins/claude-ops/skills/audit-native-overlap/scripts/overlap.py
#	plugins/claude-ops/skills/audit-native-overlap/scripts/test_overlap.py
#	plugins/claude-ops/skills/inventory/scripts/docs_crosscheck.py
#	plugins/claude-ops/skills/inventory/scripts/test_inventory.py
#	plugins/debugging/CHANGELOG.md
#	plugins/implementation/.claude-plugin/plugin.json
#	plugins/implementation/CHANGELOG.md
#	plugins/planning/.claude-plugin/plugin.json
#	plugins/planning/CHANGELOG.md
#	plugins/session-flow/.claude-plugin/plugin.json
#	plugins/session-flow/CHANGELOG.md
#	plugins/source-control/.claude-plugin/plugin.json
#	plugins/source-control/CHANGELOG.md
#	plugins/verification/CHANGELOG.md
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 29, 2026 22:06
@kyle-sexton
kyle-sexton enabled auto-merge (squash) September 29, 2026 22:06
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T22:08:57.426412Z 2584c79 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2584c7943c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/source-control/skills/pull-request/SKILL.md Outdated
…ongside /autofix-pr

Codex review on #5387: fetching alone leaves the checked-out branch on the
old tip, so a local fix commit made while /autofix-pr also pushes is based
on stale history and its push is rejected as non-fast-forward.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

…erdicts-2.1.284

# Conflicts:
#	plugins/claude-config/.claude-plugin/plugin.json
#	plugins/claude-config/CHANGELOG.md
#	plugins/claude-ops/CHANGELOG.md
#	plugins/planning/.claude-plugin/plugin.json
#	plugins/planning/CHANGELOG.md
@kyle-sexton
kyle-sexton enabled auto-merge (squash) September 29, 2026 23:07
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 1s —— View job


I'll analyze this and get back to you.

@kyle-sexton
kyle-sexton merged commit 15058ac into main Sep 29, 2026
17 of 19 checks passed
@kyle-sexton
kyle-sexton deleted the feat/native-surface-verdicts-2.1.284 branch September 29, 2026 23:11
kyle-sexton added a commit that referenced this pull request Sep 29, 2026
…earning

Main now carries the squash merges of #5371 and #5387. Their files take
main's final form; this branch re-applies only its own hunks on top.

- overlap.py: the extracted build_native_index keeps main's plugin-backed
  reclassification; detect and dismiss both call it.
- test_overlap.py: main's PluginBackedSurfaceTests kept beside the
  dismissal tests.
- records.json: main's 47 rows plus this branch's 11 rows and 58
  dismissals. Two component fingerprints recomputed where main changed
  the description (architecture:map-context, claude-ops:audit-install-state);
  docs/native-surfaces.md regenerated.
- Versions above main: claude-ops 0.67.0, claude-config 0.53.2,
  source-control 0.62.29, native-references 3.2.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…e every candidate (#5466)

No related issue: operator decisions Q8, Q9 and Q12 from the 2026-09-29
native-surfaces interview (follow-up to #5387).

## Summary

Overlap discovery re-proposed the same false positives on every run,
because a ruling of "not an overlap" had nowhere to live. About 59% of
candidates were noise. 69 discovered candidates were also still unruled,
and the new inventory and detect behavior had no evals.

Stacked on #5387, which is stacked on #5371.

## Fix

- **Dismissals (Q8):**
- The store gains an optional `dismissals` list. Each entry records the
native surface, the component, a reason, `as_of` and date, plus a
fingerprint of each side's whitespace-collapsed description.
- New `overlap.py dismiss` subcommand. It refuses a pair that already
has a verdict row.
- `detect` suppresses a dismissed pair until either fingerprint changes,
then resurfaces it flagged `resurfaced: description changed`. A verdict
row always wins.
- `self-check` validates dismissals, and `generate` renders a Dismissed
table.
- **Triage (Q9):** all 69 remaining candidates are ruled.
- 11 verdict rows, 8 of them non-defer, each with a Boundary section, or
a registry row only for agents.
  - 58 dismissals.
  - Fresh detect: 0 new candidates, 58 suppressed, 0 resurfaced.
- Every reason ends "Ruled 2026-09-29 by operator direction on the
orchestrator's recommendation."
- **Evals (Q12):**
- inventory: `is-foo-real-under-a-degraded-lane`,
`docs-crosscheck-classifies-a-removed-command`
- audit-native-overlap: `user-only-native-recommends-suggest`,
`dismissed-pair-suppressed-then-resurfaced`
- **Version bumps:** claude-ops, source-control, bugs, github,
claude-memory, claude-config, each with a CHANGELOG entry.

## Verification

- `overlap.py generate --check`: in sync (58 rows, 58 dismissals).
- `overlap.py self-check`: degraded on the 2 documented advisories only.
- `test_overlap.py`: 156 tests OK (22 new). `overlap.test.sh`: exit 0.
`test_inventory.py`: 101 tests OK.
- Eval files validate against
`plugins/skill-quality/reference/evals.schema.json`, and
`check-evals-quality.sh` passes with 0 warnings. No model evals were
run.
- `check-changed-skills.sh main`: 25 skills, 0 failed.
- `check-changelog-parity.sh`: `--check`, `--check-order` and
`--check-preserved origin/main` pass. `--check-bump origin/main` is red
only on five plugins inherited from #5387, whose versions main has since
passed; they are renumbered when main merges into the stack.
- Pinned ruff: `check` and `format --check` clean.

## Related

- #5387 is the base, and #5371 below it.
- #5465 files future drift as work items; dismissals here are what keep
that intake from repeating false positives.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
No related issue: operator decisions Q5 and Q11 from the 2026-09-29
native-surfaces interview (follow-up to #5371).

## Summary

Nothing re-ran the inventory or the overlap detector when Claude Code
shipped a release. Drift was found only by hand: new surfaces, renamed
or removed ones, invocability changes, and decision rows whose recheck
trigger had fired.

Stacked on #5371; merge that first.

## Fix

- **`/claude-ops:changelog apply` Phase 7, native-surface drift**
(`context/native-drift.md`) runs:
  - `inventory.py --self-check`
  - a full `--binary-only --docs` extraction
  - `overlap.py detect` and `overlap.py self-check`
- `native_drift.py summarize`, then `diff` against the last good
summary. A broken extraction never becomes the baseline.
- **Report:** surfaces added, removed, renamed (by alias or description
similarity) and reclassified; invocability and marker changes; docs
cross-check changes; new overlap candidates; fired store triggers.
- **Work items are filed through `/work-items:track add`** (raw intake,
`needs-triage`), one per:
  - new candidate with no store row;
  - store row whose trigger fired;
- revalidation proposal: the CLI is past `VALIDATED_AGAINST` with every
lane ok and no surface change;
  - degraded or broken inventory.
- **Dedupe:** each item carries a
`native-drift:<kind>:<surface>:<component>` key. An open item with the
key is skipped, and a candidate closed as not planned counts as
dismissed.
- **Approval:** interactive runs confirm the batch once; unattended
runs, declared by the caller, file directly.
- **Dynamic:** no surface names are hard-coded.
- `claude-ops` 0.66.0.

## Verification

- `native_drift.test.sh`: 23 cases OK.
- `changelog-status.test.sh`: 70/70.
- `overlap.test.sh`: 134 OK. `test_inventory.py`: OK.
- `check-changed-skills.sh origin/main`: 5 skills, 0 failed. The three
warnings are on lines this PR does not touch.
- `check-changelog-parity.sh`: `--check`, `--check-order`, `--check-bump
origin/main` and `--check-preserved origin/main` all pass.
- Pinned ruff: `check` and `format --check` clean. markdownlint: 0
issues.
- **Live run on 2.1.285** (validated 2.1.284, every lane ok, identical
surfaces): the revalidate path. It also flagged six store rows whose
markers moved. Four are already corrected in #5387; the two `design`
rows move to `suggest` in the planned description sweep.

## Related

- #5371 is the base. #5387 corrects the four marker rows.
- The `native-drift` label is managed as code in github-iac; a follow-up
adds it. Until then items file without it, and dedupe does not depend on
the label.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
…request descriptions (#5503)

No related issue: sweep unit 1 of the native-reference
description-phrase sweep approved in the 2026-09-29 native-surfaces
interview (Q4).

## Summary

The native-surface store records four `route` verdicts for
source-control whose routing clause lived only in each skill's Boundary
section. A body loads only on invocation, so the model picking between
our skill and the native one never saw it. A skill's description is what
the model reads when choosing a skill.

## Fix

- **`commit`:** a front-loaded, presence-gated clause for the bundled
`commit` skill and the built-in `/commit-push-pr` command.
- **`pull-request`:** a front-loaded, presence-gated clause for the
bundled `pr` skill and `/commit-push-pr`.
- **Store:** the four rows' `baked.description_phrase` flags are set,
and `docs/native-surfaces.md` is regenerated.
- **Adopters table:** the native-references Adopters table no longer
says source-control has no phrase.
- source-control 0.62.30.

Store rows gated by this change, each `complementary`, `route`, observed
by extraction against Claude Code 2.1.284:

| Native surface | Component |
|---|---|
| bundled `commit` | source-control:commit |
| built-in `/commit-push-pr` | source-control:commit |
| bundled `pr` | source-control:pull-request |
| built-in `/commit-push-pr` | source-control:pull-request |

**Not baked:** the `/autofix-pr` rows. They are `suggest` rows on a
surface the model cannot invoke. The operator ruled they get no phrase,
and the convention forbids one on that combination; their Boundary
sections already offer the command to the person.

## Verification

- `overlap.py self-check`: degraded, on the 2 documented advisories
only. `generate --check`: in sync.
- `test_overlap.py`: 172 tests OK.
- `check-changed-skills.sh origin/main`: 0 failed.
- `validate-plugin-contracts.mjs`: 0 warnings.
- `check-changelog-parity.sh`: all four modes pass.
- `check-spoke-plugin-root.sh`, typos and the ai-slop report: all clean.
- Description lengths: 654/1536 for `commit`, 659/1536 for
`pull-request`.

## Related

- #5387 recorded these verdicts and their Boundary sections.
- Sweep contract: one plugin per PR, each merged before the next unit
starts (`audit-native-overlap/SKILL.md`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant